As we all know in India; agriculture is the country\'s backbone. This report forecasts the yield of practically every type of crop grown in India. This programme is created by utilising simple characteristics such as state, district, and the agricultural yield can be predicted based on the season, area, and the user. the year he or she wishes to participate in. The paper makes use of machine learning algorithms like linear regression and random forest. I have also discussed why predicting on the basis of the time series algorithm is a bad idea. And with the multiple factors in the picture like average rain, political concerns, current economy, using time series models will be a bad idea since the predictions will be of no use or with minimum accuracy.
Introduction
This study examines India's agricultural sector and explores the use of machine learning techniques to predict crop production. Although agriculture's contribution to India's GDP has declined over the years due to growth in other sectors, it remains vital to the economy, supporting around 70% of the workforce, providing food security, supplying raw materials to industries, and contributing significantly to exports.
The research uses agricultural data obtained from the Government of India's official open data portal, containing 246,091 records from 1997 to 2015. The dataset includes information on states, districts, crop year, season, crop type, cultivation area, and production. After data collection, the dataset is cleaned, preprocessed, and analysed using exploratory data analysis (EDA) to identify trends before selecting suitable machine learning models.
The exploratory analysis reveals regional differences in agricultural production across India. The South Zone records the highest overall production, largely driven by Kerala's coconut cultivation. Variations in production around 2005 and 2011 were linked to differences in monsoon rainfall patterns, highlighting the strong influence of climatic conditions on agricultural output.
A survey of 64 respondents, including farmers, students, researchers, agribusiness professionals, and agricultural officers, showed that most farmers own small land holdings (1–5 acres). Kharif was identified as the most important cropping season, while rice was the most commonly cultivated crop. Most respondents rated their crop yields as good, and 77.8% reported awareness of digital agricultural technologies such as weather forecasting, IoT-based monitoring, and AI-driven decision-support systems.
The study evaluated time series forecasting but found it unsuitable because agricultural production lacked stable temporal trends and was heavily influenced by external factors such as rainfall and weather. Since historical production data alone could not reliably predict future yields, the researchers proposed using supervised machine learning methods, particularly Linear Regression, which can incorporate multiple explanatory variables and better capture the complex relationships affecting agricultural production. The study concludes that machine learning approaches are more appropriate than traditional time series models for predicting crop production under changing environmental conditions.
Conclusion
The precise prediction of agricultural output continues to pose a significant challenge, primarily due to the considerable impact of uncertain and dynamic elements, including weather patterns, climate fluctuations, and local farming practices. This research underscored the critical role of exploratory data analysis in the preparation and comprehension of agricultural datasets, as it aids in data cleansing, trend detection, and pattern recognition. By utilizing agricultural production data from India, the effectiveness of supervised machine learning models was assessed, with Random Forest identified as the most proficient method. This model adeptly captured intricate, non-linear interactions among the variables and demonstrated enhanced predictive accuracy when compared to conventional techniques such as Linear Regression. The results emphasize the promise of machine learning methodologies in improving agricultural forecasting and facilitating evidence-based decision-making. Additionally, the reliance on publicly accessible government datasets promotes transparency and reproducibility in the research process.
The suggested methodology can support policymakers, agricultural organizations, and farmers in strategic planning and resource distribution, ultimately leading to enhanced agricultural productivity, sustainability, and food security.
References
[1] Chand, R., Raju, S. S., & Pandey, L. M. (2017). Growth crisis in Indian agriculture: Severity and options at national and state levels. Economic & Political Weekly, 52(21).
[2] Chlingaryan, A., Sukkarieh, S., & Whelan, B. (2018). Machine learning approaches for crop yield prediction and nitrogen status estimation. Computers and Electronics in Agriculture, 151, 61–69.
[3] Birthal, P. S., Joshi, P. K., Roy, D., & Thorat, A. (2012). Diversification in Indian agriculture towards high-value crops. Agricultural Economics Research Review, 25(1), 1–12.
[4] Jeong, J. H., Resop, J. P., Mueller, N. D., et al. (2016). Random forests for global and regional crop yield predictions. PLOS ONE, 11(6).
[5] Ministry of Agriculture & Farmers Welfare. State of Indian Agriculture. Government of India, New Delhi.
[6] National Sample Survey Office (NSSO). Situation Assessment Survey of Agricultural Households in India. Ministry of Statistics and Programme Implementation, Government of India